Linear Mixed Model
First, What does mixed-effects mean
Mixed-effects models
- A mixed-effects model is a statistical model that mixes both:
- fixed effects: population-level effects (effects we want to estimate directly)
- random effects: group-level or individual-level deviations from the population-average effects
- Useful when data are not fully independent, but the dependency has a meaningful structure
- example: repeated measurements from the same subject (Longitudinal Data Analysis)
- example: students nested within schools
- example: patients nested within hospitals
- The core idea: model the average pattern while also accounting for individual / group differences
| Model | Emphasize |
| ---------------------- | -------------------------------------- |
| Mixed model | fixed + random effects |
| Multilevel model | nested / clustered data structure |
| Hierarchical model | hierarchical model/parameter structure |
- They all describe the same framework from different angles.
- Hierarchical model is the broadest term.
Fixed effects vs. random effects
Fixed effects
- Fixed effects answer: what is the average relationship in the population?
- Example:
- if we study the effect of treatment, the treatment effect is usually a fixed effect
- it represents the average treatment effect across the whole population
Random effects
- Random effects answer: how much do individuals or groups deviate from that average?
- Example:
- each subject may have a different baseline level / average performance
- each patient may respond differently to time
Linear mixed model (LMM)
What is LMM
- Linear mixed model (LMM), also called linear mixed-effects model:
- a type of mixed-effects model for continuous outcomes (but not all mixed-effects models are LMMs; see Generalized Linear Model)
- "Linear": the response is modeled as a linear combination of predictors, similar to Regression
- LMM is useful when the data have both:
- within-group variability: e.g., repeated measurements from the same subject
- between-group variability: e.g., differences between subjects, hospitals, schools, batches, etc.
- LMM is mainly for continuous outcomes that are approximately normally distributed.
- For binary, count, ordinal, or categorical outcomes, use a generalized linear mixed model (GLMM) instead.
Formula
: response vector : design matrix for fixed effects : fixed-effect coefficients : design matrix for random effects : random effects : residual error
Intuition of the formula
= population-level average pattern = subject-level / group-level deviation from the average pattern = leftover noise not explained by the model
Example: Random Intercept + Random Slope
For Longitudinal Data Analysis, a common LMM is:
: population-average intercept : population-average rate of change : subject-specific deviation from the average intercept : subject-specific deviation from the average slope
Thus, each individual has their own trajectory:
Variance components
Random effects describe how individuals differ from the population-average. Their variability is summarized by variance components.
Variance components represent remaining unexplained variability after accounting for the fixed effects.
- Large random-intercept variance → individuals still differ in baseline levels.
- Large random-slope variance → individuals still differ in rates of change.
- Large (level-1/within-person) residual variance → substantial within-person variation remains unexplained.
- Intercept–slope covariance:
- positive → individuals with higher-than-average baselines tend to have higher slopes
- negative → individuals with higher-than-average baselines tend to have lower slopes
- near zero → little association between baseline deviation and rate-of-change deviation
They helps identify where additional predictors may be useful:
- within-person variation → consider time-varying predictors
- between-person variation → consider individual/group-level predictors
Example: Random Intercept + Random Slope
For a random-intercept + random-slope model:
where
: between-person variability in intercepts : between-person variability in rates of change : association between intercept and rate of change : within-person residual variability around each individual's trajectory
Why LMM matters
Compared to ANOVA
Compared to ANOVA & Post-hoc Tests, LMM is more flexible:
- It can handle non-independent observations, such as repeated measures (Longitudinal Data Analysis) or clustered data
- It can model correlation within the same subject or group
- It can handle unbalanced data, where different groups have different numbers of observations
- It can often handle missing observations better than repeated-measures ANOVA, as long as the missingness assumption is reasonable
When to use LMM
Use LMM when:
- the outcome is continuous
- observations are correlated within subject / group / cluster
- each subject or group can have its own baseline level or slope
- the data are hierarchical, nested, longitudinal, or repeated-measures
Assumptions for using LMM
Sampling and independence
- Subjects or groups are sampled from the population of interest.
- Observations from different individuals / groups are independent.
- Repeated measurements from the same individual are not assumed to be independent; this dependence is modeled through random effects and/or covariance structure.
Distribution assumptions
- Conditional on the fixed and random effects, residuals are approximately normally distributed.
- Random effects are usually assumed to be normally distributed.
- Missing data are assumed to be ignorable, often interpreted as missing at random (MAR).
Covariance structure
How to build an LMM
General modeling decisions:
- Decide fixed effects
- What population-level relationships should be estimated?
- Decide random effects
- Which grouping variable should have random effects?
- Random intercept only, or random intercept + random slope?
- Decide Covariance Structure
- Which structure best describes within-subject or within-group correlation?
- Compare candidate models
- Use likelihood ratio tests, AIC / BIC, or cross-validation
- Prefer the model that is interpretable and fits the dependency structure well
- Avoid making the random-effects structure too complex if the data cannot support it
For repeated-measures / growth-curve applications, see Longitudinal Mixed Model Workflow.